Efficient Utilization of Rare Variants for Detection of Disease-Related Genomic Regions
نویسندگان
چکیده
When testing association between rare variants and diseases, an efficient analytical approach involves considering a set of variants in a genomic region as the unit of analysis. One factor complicating this approach is that the vast majority of rare variants in practical applications are believed to represent background neutral variation. As a result, analyzing a single set with all variants may not represent a powerful approach. Here, we propose two alternative strategies. In the first, we analyze the subsets of rare variants exhaustively. In the second, we categorize variants selectively into two subsets: one in which variants are overrepresented in cases, and the other in which variants are overrepresented in controls. When the proportion of neutral variants is moderate to large we show, by simulations, that the both proposed strategies improve the statistical power over methods analyzing a single set with total variants. When applied to a real sequencing association study, the proposed methods consistently produce smaller p-values than their competitors. When applied to another real sequencing dataset to study the difference of rare allele distributions between ethnic populations, the proposed methods detect the overrepresentation of variants between the CHB (Chinese Han in Beijing) and YRI (Yoruba people of Ibadan) populations with small p-values. Additional analyses suggest that there is no difference between the CHB and CHD (Chinese Han in Denver) datasets, as expected. Finally, when applied to the CHB and JPT (Japanese people in Tokyo) populations, existing methods fail to detect any difference, while it is detected by the proposed methods in several regions.
منابع مشابه
Detection of Genetic Differences between Holstein and Iranian North-West Indigenous Hybrid Cattles using Genomic Data
Extended Abstract Introduction and Objective: Selection to increase the frequency of new mutations useful only in some subpopulations leaves markers at the genome level. Most of these regions are related to genes and QTLs controlling significant economic traits. Material and Methods: In order to detection of genetic differences between Iranian northwestern crossbred and Holstein cattle breed,...
متن کاملSurvey on Perception of People Regarding Utilization of Computer Science & Information Technology in Manipulation of Big Data, Disease Detection & Drug Discovery
this research explores the manipulation of biomedical big data and diseases detection using automated computing mechanisms. As efficient and cost effective way to discover disease and drug is important for a society so computer aided automated system is a must. This paper aims to understand the importance of computer aided automated system among the people. The analysis result from collected da...
متن کاملA Simple Genome Walking Strategy to Isolate Unknown Genomic Regions Using Long Primer and RAPD Primer
Background: Genome walking is a DNA-cloning methodology that is used to isolate unknown genomic regions adjacent to known sequences. However, the existing genome-walking methods have their own limitations. Objectives: Our aim was to provide a simple and efficient genome-walking technology. Material and Methods: In this paper, we dev...
متن کاملCRB1-Related Leber Congenital Amaurosis: Reporting Novel Pathogenic Variants and a Brief Review on Mutations Spectrum
Background: Leber congenital amaurosis (LCA) is a rare inherited retinal disease causing severe visual impairment in infancy. It has been reported that 9-15% of LCA cases have mutations in CRB1 gene. The complex of CRB1 protein with other associated proteins affects the determination of cell polarity, orientation, and morphogenesis of photoreceptors. Here, we report three novel pathogenic varia...
متن کاملRare-Allele Detection Using Compressed Se(que)nsing
Detection of rare variants by resequencing is important for the identification of individuals carrying disease variants. Rapid sequencing by new technologies enables low-cost resequencing of target regions, although it is still prohibitive to test more than a few individuals. In order to improve cost trade-offs, it has recently been suggested to apply pooling designs which enable the detection ...
متن کامل